feat(vllm): auto-prepare trusted dual DGX Stations - #7030
Conversation
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Senthil Kumar Ravichandran <senthilr@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
📝 WalkthroughWalkthroughThis PR adds DGX Station dual-peer qualification, secure SSH and resume state, managed two-node vLLM lifecycle orchestration, onboarding integration, updated documentation, and a guarded fixture-backed simulator with extensive tests. ChangesDual-Station installer and preparation
Estimated code review effort: 5 (Critical) | ~120 minutes Possibly related PRs
Suggested labels: Suggested reviewers: 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
Code Coverage OverviewLanguages: TypeScript TypeScript / code-coverage/pluginThe overall coverage in commit f82e692 in the TypeScript / code-coverage/cliThe overall coverage in commit f82e692 in the Show a code coverage summary of the most impacted files.
Updated |
PR Review Advisor — No blocking findings reportedAdvisor assessment: No blocking advisor findings reported Model lanes
Nemotron output stays in workflow artifacts and does not change the assessment above. E2E guidanceAdvisory only. E2E / PR Gate selects and runs jobs independently. Recommended E2E: This automated review informs maintainers. Warnings and suggestions do not require a response. A maintainer decides whether to merge. |
Signed-off-by: Senthil Kumar Ravichandran <senthilr@nvidia.com>
…st-prereqs Signed-off-by: Senthil Kumar Ravichandran <senthilr@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com> (cherry picked from commit 586a60d)
Signed-off-by: Senthil Ravichandran <senthilr@nvidia.com>
Signed-off-by: Senthil Ravichandran <senthilr@nvidia.com>
Signed-off-by: Senthil Ravichandran <senthilr@nvidia.com>
Signed-off-by: Senthil Ravichandran <senthilr@nvidia.com>
Signed-off-by: Senthil Ravichandran <senthilr@nvidia.com>
Signed-off-by: Senthil Ravichandran <senthilr@nvidia.com>
Signed-off-by: Senthil Ravichandran <senthilr@nvidia.com>
Signed-off-by: Senthil Ravichandran <senthilr@nvidia.com>
Signed-off-by: Senthil Ravichandran <senthilr@nvidia.com>
|
@prekshivyas The physical and trusted-boundary blockers from the review of Physical dual-Station resultTwo physical DGX Station GB300 systems on the accepted direct private
Timings:
All observed GB300 and auxiliary RTX PRO volatile ECC counters ended at corrected/uncorrected The cold product run used The implementation and documentation remain within the accepted Deferred boundary: exactly two trusted Stations, direct private Please re-review or dismiss the stale changes-requested state after the current CI run settles. |
Summary
Station Express detects one pretrusted DGX Station GB300 peer across two direct private
/30rails and prepares the qualified pair for distributed Nemotron 3 Ultra inference. Without a qualified pair, Express keeps the existing single-Station Ultra path. The dual-Station path remains Deferred and does not change the single-Station support status. Operators still own rail configuration, network isolation, SSH trust, firewall policy, and reboots.Changes
/30, 400 Gb/s, MTU-9000 CX-8 rails. The installer does not scan subnets or enroll SSH trust.NEMOCLAW_DGX_STATION_PEER,NEMOCLAW_VLLM_MODEL, and--station-deepseekauthoritative and fail closed on conflicting intent.scripts/prepare-dgx-station-host.shon both hosts with exact-byte hashing and strict noninteractive SSH execution.nemotron-ultra. Single-Station Ultra continues to servenvidia/nemotron-3-ultra-550b-a55b.O_NOFOLLOW, and only explicit managed bindings enter managed lifecycle handling.--checkto use sudo only for read-only Docker inventory, reopening SSH after peer login-required status, and binding both controller UIDs after pair qualification.HOMEhas a trailing slash and accept physicaliproute2route JSON that omits redundant filtered fields while retaining exact direct-route, device, scope, MAC, neighbor-state, and jumbo-frame checks.Type of Change
Quality Gates
Documentation Writer Review
passmain; its readiness and E2E-support changes do not alter dual-Station behavior, and the seven reviewed dual-Station documentation files are byte-identical to the prior reviewed head. Installer integration passed521tests with2intentional skips, focused CLI passed103/103, E2E support passed19/19, the integration workflow test passed3/3, CLI typecheck passed, and the diff check passed.DGX Station Hardware Evidence
4615f2523d179da6d872fe33b1d226804b36c58f; cold-install headdd03ed8e45fb0e5f1e1ffa7f1fdeb65d5e6a0a94. Current headf82e69259dbf124491d645abb0d73d7cf8214704adds only documentation, current-mainintegration, CI fixture/module-reference changes, readiness host-observation reporting, and E2E restore-result classification after the behavior head./30rails. Product-default OpenClaw + Nemotron 3 Ultra completed reciprocal preparation, authenticated two-node Ray/vLLM serving, chat/tool smoke, managed-runtime restart recovery, warm reuse, durable cleanup ownership, and exact-pair rollback. The worker's auxiliary RTX PRO remained excluded from managed inference. All observed volatile ECC counters ended at corrected/uncorrected0/0.24m06.994s; chat completed in5.57s; tool smoke completed in23.50s; managed-runtime restart recovery completed in9m50s; exact warm reuse completed in2m51.300s; the post-reuse prompt completed in7.57s; full rollback completed in10.900s. Rollback removed the exact pair and receipt while preserving the image, model caches, Docker, Toolkit, OpenShell, controller binding, driver, and rail configuration. No credential was persisted or included in evidence.Verification
Signed-off-by:line and every commit appears asVerifiedin GitHubpre-commit,commit-msg, andpre-pushhooks passed, ornpm run check:diffpassed when hooks were skipped or unavailable521tests with2intentional skips; focused dual-runtime CLI tests passed83/83; CLI typecheck, Biome, shell syntax, diff checks, and hooks passed.f82e69259.npm run docsbuilds without warnings (doc changes only) — build passed with 0 errors and 2 existing Fern warnings.Signed-off-by: Aaron Erickson aerickson@nvidia.com
Signed-off-by: Senthil Ravichandran senthilr@nvidia.com